iT邦幫忙

2026 iThome 鐵人賽

DAY 11
0

前言

Day 10 我們建立了第一組 Eval Dataset。

目前 evals/cases.json 裡已經有 15 筆測試案例,包含:

  • calculation
  • keyword_qa
  • general_qa
  • instruction_following
  • json_output

但到目前為止,這些案例還只是靜態資料。

如果要測 Agent,我們還是得手動複製每一題,貼到 python3 app.py 裡執行。

這樣不適合做 Evaluation。

所以 Day 11 開始讓測試流程自動化:

讀取 evals/cases.json,逐筆呼叫 Agent,儲存每題的 Agent output,並綁定對應的 trace session id。

今天還不會判斷 pass / fail。

今天先讓整批測試案例跑起來。


今天要完成什麼?

會完成:

  1. 讀取 eval dataset。
  2. 逐筆呼叫 Agent。
  3. 儲存每題 Agent output。
  4. 綁定 trace session id。
  5. 將每次 batch run 的結果輸出成 JSON。
  6. 在終端機印出簡單測試結果表。

今天先不做:

  • 自動評分。
  • success rate。
  • failure type。
  • JSON schema validation。
  • Evaluation dashboard。

這些從 Day 12 開始逐步處理。


Batch Evaluation Runner 是什麼?

Batch Evaluation Runner 可以想成一個批次測試器。

它負責把 eval dataset 裡的測試案例一題一題丟給 Agent。

流程如下:

讀取 cases.json
  -> for each test case
  -> 呼叫 agent.run(input)
  -> 取得 AgentResult
  -> 保存 trace
  -> 記錄 output 和 session_id
  -> 輸出 evaluation run 結果

今天的 runner 還不會判斷答案是否正確。

所以它不會輸出:

case_001: pass
case_002: fail

而是先輸出:

case_001: completed, session_id=...
case_002: completed, session_id=...

這樣拆,是因為 Evaluation 可以分成兩層。

第一層是執行:

把所有 test cases 跑完,收集 Agent output。

第二層是評分:

根據 expected 和 grading_method 判斷 pass / fail。

Day 11 先處理第一層,Day 12 再處理第二層。


今天的專案結構

今天新增一個 runner 檔案,並讓 batch run 結果輸出到 data/eval_runs/。

agent-testing-platform/
  app.py
  trace_viewer_app.py
  agents/
    __init__.py
    simple_agent.py
    prompts.py
    fake_llm.py
  tools/
    __init__.py
    calculator.py
  tracing/
    __init__.py
    models.py
  storage/
    __init__.py
    database.py
    schema.sql
  ui/
    __init__.py
    trace_viewer.py
  evals/
    __init__.py
    cases.json
    runner.py
  data/
    agent_traces.db
    eval_runs/
      eval_run_*.json

新增:

檔案 用途
evals/runner.py 讀取 eval dataset,逐筆執行 Agent,輸出 batch run 結果

今天不修改:

  • agents/simple_agent.py
  • storage/database.py
  • evals/cases.json

因為 Day 11 的重點是新增 batch runner,不是修改 Agent 行為或測試資料。


設計 Batch Run 結果格式

執行完一整批 test cases 後,我們會產生一份 JSON 結果。

格式大概長這樣:

{
  "run_id": "eval_run_20260906_103000",
  "total_cases": 15,
  "results": [
    {
      "case_id": "case_001",
      "input": "請計算 135 * 28",
      "expected": "3780",
      "grading_method": "contains",
      "task_type": "calculation",
      "status": "completed",
      "actual": "The result is 3780",
      "trace_session_id": "..."
    }
  ]
}

今天先記錄這些欄位:

欄位 說明
run_id 這次 batch run 的編號
total_cases 總共跑了幾筆案例
case_id 對應原本的 test case
input Agent 收到的任務
expected 預期答案,今天先保存但不評分
grading_method 評分方式,今天先保存但不使用
task_type 任務類型
status 目前只分成 completed 或 error
actual Agent 實際輸出
trace_session_id 對應到 Trace Viewer 的 session id
error 如果執行失敗,記錄錯誤訊息

trace_session_id 很有用。

後面看到某一題失敗時,可以用這個 id 回 Trace Viewer 查看當時 Agent 中間做了什麼。


實作 Batch Evaluation Runner

新增 evals/runner.py:

import json
from datetime import datetime
from pathlib import Path
from typing import Any

from agents.fake_llm import FakeLLMClient
from agents.simple_agent import SimpleAgent
from storage.database import init_db, save_trace


BASE_DIR = Path(__file__).resolve().parent.parent
CASES_PATH = BASE_DIR / "evals" / "cases.json"
EVAL_RUNS_DIR = BASE_DIR / "data" / "eval_runs"


def load_cases() -> list[dict[str, Any]]:
    with CASES_PATH.open(encoding="utf-8") as file:
        return json.load(file)


def create_run_id() -> str:
    timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
    return f"eval_run_{timestamp}"


def run_evaluation() -> dict[str, Any]:
    init_db()

    agent = SimpleAgent(llm_client=FakeLLMClient())
    cases = load_cases()
    run_id = create_run_id()
    results = []

    for test_case in cases:
        try:
            result = agent.run(test_case["input"])
            save_trace(result.trace)

            results.append(
                {
                    "case_id": test_case["id"],
                    "input": test_case["input"],
                    "expected": test_case["expected"],
                    "grading_method": test_case["grading_method"],
                    "task_type": test_case["task_type"],
                    "status": "completed",
                    "actual": result.answer,
                    "trace_session_id": result.trace.session_id,
                    "error": None,
                }
            )
        except Exception as exc:
            results.append(
                {
                    "case_id": test_case["id"],
                    "input": test_case["input"],
                    "expected": test_case["expected"],
                    "grading_method": test_case["grading_method"],
                    "task_type": test_case["task_type"],
                    "status": "error",
                    "actual": None,
                    "trace_session_id": None,
                    "error": str(exc),
                }
            )

    eval_run = {
        "run_id": run_id,
        "created_at": datetime.now().isoformat(),
        "total_cases": len(cases),
        "results": results,
    }

    save_eval_run(eval_run)
    return eval_run


def save_eval_run(eval_run: dict[str, Any]) -> Path:
    EVAL_RUNS_DIR.mkdir(parents=True, exist_ok=True)
    output_path = EVAL_RUNS_DIR / f"{eval_run['run_id']}.json"

    with output_path.open("w", encoding="utf-8") as file:
        json.dump(eval_run, file, ensure_ascii=False, indent=2)

    return output_path


def print_summary(eval_run: dict[str, Any]) -> None:
    print(f"Run ID: {eval_run['run_id']}")
    print(f"Total cases: {eval_run['total_cases']}")
    print()

    for result in eval_run["results"]:
        print(
            f"{result['case_id']} | "
            f"{result['task_type']} | "
            f"{result['status']} | "
            f"trace={result['trace_session_id']}"
        )


if __name__ == "__main__":
    eval_run = run_evaluation()
    print_summary(eval_run)

這支檔案就是今天的 Batch Evaluation Runner。

它主要分成幾個 function。


load_cases:讀取測試資料

load_cases() 會讀取 evals/cases.json:

def load_cases() -> list[dict[str, Any]]:
    with CASES_PATH.open(encoding="utf-8") as file:
        return json.load(file)

這裡使用 encoding="utf-8",是因為 cases.json 裡有中文任務。

如果沒有指定編碼,在某些環境中可能會遇到中文讀取問題。


run_evaluation:逐筆執行 Agent

run_evaluation() 是今天最主要的 function。

它會先初始化資料庫:

init_db()

接著建立 Agent:

agent = SimpleAgent(llm_client=FakeLLMClient())

然後讀取所有 cases:

cases = load_cases()

接著逐筆執行:

for test_case in cases:
    result = agent.run(test_case["input"])
    save_trace(result.trace)

這裡有一個重要設計:

每一筆 test case 執行後,都會把 trace 存進 SQLite。

所以 batch run 的結果會有 trace_session_id,可以對應回 Trace Viewer。


為什麼要記錄 trace_session_id?

假設 Day 12 之後開始評分,我們看到:

case_013 failed

如果沒有 trace,我們只知道它失敗。

但如果 evaluation result 裡有:

{
  "case_id": "case_013",
  "trace_session_id": "34ddf4d0-19c3-45d1-b44f-2d5fd6d9b42a"
}

就可以去 Trace Viewer 選這個 session,查看:

  • user input 是什麼。
  • LLM response 是什麼。
  • 有沒有呼叫工具。
  • tool input 是什麼。
  • tool output 是什麼。
  • final answer 是什麼。

這就是 Trace 和 Eval 串起來的地方。

Eval 告訴我們哪一題有問題。

Trace 幫助我們看出問題發生在哪一步。


錯誤案例先記錄 error,不做分類

在 run_evaluation() 裡,我們用 try / except 包住每一筆 test case:

try:
    result = agent.run(test_case["input"])
    save_trace(result.trace)
    ...
except Exception as exc:
    results.append(
        {
            "status": "error",
            "actual": None,
            "trace_session_id": None,
            "error": str(exc),
        }
    )

這樣做是為了避免某一題失敗時,整個 batch run 中斷。

例如某一題讓 Agent 丟出 exception,runner 還是可以繼續跑下一題。

今天只先記錄:

status: error
error: 錯誤訊息

至於這個錯誤是 format_error、tool_error 還是 instruction_error,會留到第三週 Failure Analysis 再處理。


save_eval_run:保存 batch run 結果

save_eval_run() 會把整次 batch run 結果輸出成 JSON:

def save_eval_run(eval_run: dict[str, Any]) -> Path:
    EVAL_RUNS_DIR.mkdir(parents=True, exist_ok=True)
    output_path = EVAL_RUNS_DIR / f"{eval_run['run_id']}.json"

    with output_path.open("w", encoding="utf-8") as file:
        json.dump(eval_run, file, ensure_ascii=False, indent=2)

    return output_path

輸出位置會在:

data/eval_runs/

檔名會類似:

eval_run_20260906_103000.json

這樣每次 batch run 都會留下紀錄。

之後 Day 12 加上 evaluator 後,就可以在同一份結果裡加入 pass / fail。


print_summary:先用終端機查看結果

今天還沒有做 dashboard,所以先用終端機印出簡單結果表。

print_summary() 會輸出:

def print_summary(eval_run: dict[str, Any]) -> None:
    print(f"Run ID: {eval_run['run_id']}")
    print(f"Total cases: {eval_run['total_cases']}")
    print()

    for result in eval_run["results"]:
        print(
            f"{result['case_id']} | "
            f"{result['task_type']} | "
            f"{result['status']} | "
            f"trace={result['trace_session_id']}"
        )

目前它只顯示:

  • case id
  • task type
  • status
  • trace session id

等 Day 12 開始自動評分後,這裡會再加入:

  • passed
  • failure reason

執行 Batch Evaluation Runner

在專案根目錄執行:

python3 -m evals.runner

預期會看到類似結果:

Run ID: eval_run_20260906_103000
Total cases: 15

case_001 | calculation | completed | trace=8a2b2f21-d1a4-4d8d-9a57-0db43a4e9d23
case_002 | calculation | completed | trace=9c75f807-3c8e-47b2-a38c-55f7b9b2161a
case_003 | calculation | completed | trace=0ad5ef35-334f-44f4-906e-2e95d534f6ac
...
case_015 | json_output | completed | trace=7d734690-0e49-4e3e-8f1e-1d98dbeb820d

如果某一題執行時發生錯誤,會看到:

case_xxx | calculation | error | trace=None

這表示 runner 沒有因為單一錯誤中斷,而是繼續執行其他 test cases。


檢查輸出的 eval run 檔案

執行完成後,可以查看 data/eval_runs/:

ls data/eval_runs

應該會看到類似:

eval_run_20260906_103000.json

也可以用 python3 -m json.tool 檢查輸出的 JSON:

python3 -m json.tool data/eval_runs/eval_run_20260906_103000.json

實際檔名要換成你當次產生的檔案名稱。

裡面會包含每一題的:

  • case_id
  • input
  • expected
  • grading_method
  • task_type
  • status
  • actual
  • trace_session_id
  • error

用 Trace Viewer 查看 batch run 產生的 trace

Day 11 的 runner 會對每一筆成功執行的 test case 呼叫:

save_trace(result.trace)

所以執行 batch run 後,Trace Viewer 裡應該會多出多筆 session。

啟動 Trace Viewer:

streamlit run trace_viewer_app.py

接著可以選擇某一筆 session,查看該 test case 的執行過程。

這表示我們已經把兩個系統接起來:

Batch Evaluation Runner
  -> Agent Runner
  -> Trace Storage
  -> Trace Viewer

雖然現在還沒有評分,但已經可以批次產生可追蹤的 Agent 執行紀錄。


為什麼今天不做 pass / fail?

今天很容易想順手把 pass / fail 做完。

例如:

passed = expected in actual

但依照這個系列的節奏,Day 11 只處理 batch execution。

原因是「執行」和「評分」是兩個不同責任。

Batch Runner 負責:

把 test cases 跑完,收集 Agent output。

Evaluator 負責:

根據 grading_method 判斷 output 是否符合 expected。

如果今天把兩者混在一起,程式會很快變大,也不利於後面擴充不同 grading method。

所以 Day 12 會專門處理 evaluator,先實作最簡單的:

  • exact match
  • keyword / contains match

今天完成後的系統狀態

今天完成後,系統具備:

  • 可以讀取 evals/cases.json。
  • 可以逐筆呼叫 SimpleAgent。
  • 可以保存每筆 test case 的 trace。
  • 可以把 batch run 結果輸出成 JSON。
  • 可以在終端機看到簡單測試結果表。
  • 可以透過 trace_session_id 回到 Trace Viewer 查看執行過程。

目前還沒有:

  • 自動評分。
  • pass / fail。
  • success rate。
  • failure reason。
  • evaluation dashboard。

今天的重點整理

Day 10 建立了 eval dataset。

Day 11 則讓 dataset 真的跑起來。

今天最重要的流程是:

cases.json
  -> evals.runner
  -> SimpleAgent.run()
  -> AgentResult
  -> save_trace()
  -> eval_run_*.json

這表示我們已經從「手動測試」前進到「批次執行」。

雖然目前還沒有判斷對錯,但這一步很重要。

因為只有先穩定收集每題的 Agent output,後面才有辦法做自動評分、成功率統計與錯誤分析。


下一步

Day 12 會實作最簡單的自動評分。

下一篇會新增 evaluator,先處理兩種 grading method:

  • exact_match
  • contains

到時候 evaluation result 就會從:

{
  "status": "completed",
  "actual": "The result is 3780"
}

進一步變成:

{
  "status": "completed",
  "actual": "The result is 3780",
  "passed": true,
  "failure_reason": null
}

這樣平台就會開始回答第二週最關心的問題:

Agent 到底有沒有完成任務?


上一篇
Day 10|建立第一組 Eval Dataset
下一篇
Day 12|實作最簡單的自動評分
系列文
從黑盒到可驗證:30 天打造 AI Agent 的 Trace、Eval 與 Guardrails 系統 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言